Papers with objective metrics
Controllable Neural Dialogue Summarization with Personal Named Entity Planning (2021.emnlp-main)
Copied to clipboard
| Challenge: | Experimental results show that our proposed framework generates fluent and factually consistent summaries under various planning controls using both objective metrics and human evaluations. |
| Approach: | They propose a controllable neural generation framework that can guide dialogue summarization with personal named entity planning. |
| Outcome: | The proposed framework generates fluent and factually consistent summaries under various planning controls using objective metrics and human evaluations. |
Tractable & Coherent Multi-Document Summarization: Discrete Optimization of Multiple Neural Modeling Streams via Integer Linear Programming (2022.emnlp-industry)
Copied to clipboard
| Challenge: | Multi-document summarization generates summary of corpus of documents consisting of related topics. |
| Approach: | They propose a generic framework to jointly consider coherence and informativeness in multi-document summarization and offers provisions to replace individual components based on the domain of source text. |
| Outcome: | The proposed framework consistently performs better than baselines for objective metrics and human evaluation. |
Coherent and Concise Radiology Report Generation via Context Specific Image Representations and Orthogonal Sentence States (2021.naacl-industry)
Copied to clipboard
| Challenge: | Neural models for text generation are often designed in an end-to-end fashion, limiting their practical usability in downstream applications. |
| Approach: | They propose a method to compute image representations specific to each sentential context and exploiting diverse sentence states to ensure topical continuity and content diversity of generated radiology reports. |
| Outcome: | The proposed method outperforms baselines on objective metrics and human evaluations by 18% and 29% respectively in the evaluation for informativeness and content ordering respectively. |
Low-Resource Multilingual and Zero-Shot Multispeaker TTS (2022.aacl-main)
Copied to clipboard
| Challenge: | Currently, the amount of data needed for TTS is limited to the vast majority of the spoken languages. |
| Approach: | They propose to use language agnostic meta learning procedure to learn speaking a new language with just 5 minutes of training data while retaining the ability to infer the voice of even unseen speakers. |
| Outcome: | The proposed approach is able to learn speaking a new language using just 5 minutes of training data while retaining the ability to infer the voice of even unseen speakers in the newly learned language. |
ETOM: A Five-Level Benchmark for Evaluating Tool Orchestration within the MCP Ecosystem (2026.findings-eacl)
Copied to clipboard
| Challenge: | Existing benchmarks assess tools in isolation, overlooking challenges such as functional overlap and cross-server orchestration, which can lead to overly optimistic evaluations. |
| Approach: | They propose a five-level benchmark for evaluating multi-hop, end-to-end tool orchestration by LLM agents within a hierarchical Model-Context Protocol (MCP) ecosystem. |
| Outcome: | The proposed framework evaluates end-to-end tool orchestration by agents in hierarchical Model-Context Protocol (MCP) environments. |
ChartMind: A Comprehensive Benchmark for Complex Real-world Multimodal Chart Question Answering (2025.emnlp-main)
Copied to clipboard
| Challenge: | Chart question answering (CQA) is a multimodal task for evaluating the reasoning capabilities of vision-language models. |
| Approach: | They propose a chart question answering benchmark that incorporates multilingual contexts and supports open-domain textual outputs. |
| Outcome: | The proposed framework outperforms the previous three common CQA paradigms: instruction-following, OCR-enhanced, and chain-of-thought. |
From Machine Translation to Code-Switching: Generating High-Quality Code-Switched Text (2021.acl-long)
Copied to clipboard
| Challenge: | a computational model for code-switching text is lacking in the corpus of real text. |
| Approach: | They propose a neural machine translation model to generate Hindi-English code-switched sentences using monolingual Hindi sentences. |
| Outcome: | The proposed model reduces perplexity on a language modeling task and improves on linguistic inference tasks. |
Leveraging the Interplay between Syntactic and Acoustic Cues for Optimizing Korean TTS Pause Formation (2024.lrec-main)
Copied to clipboard
| Challenge: | despite recent advances in speech synthesis, the focus of research has been on high-resource languages like English. |
| Approach: | They propose a framework that incorporates modeling of syntactic and acoustic cues associated with pausing patterns. |
| Outcome: | The proposed framework generates natural speech even for longer and intricate out-of-domain sentences, despite training on short audio clips. |
InspireDebate: Multi-Dimensional Subjective-Objective Evaluation-Guided Reasoning and Optimization for Debating (2025.acl-long)
Copied to clipboard
| Challenge: | Existing LLMs focus on responding to specific arguments while neglecting objective assessments such as authenticity and logical validity. |
| Approach: | They propose a multi-dimensional evaluation system and an optimized debating framework . they propose to use coT reasoning enhancement, web-based Retrieval Augmented Generation to optimize across various dimensions. |
| Outcome: | The proposed framework outperforms baseline models in argument quality assessment and debate process simulation by 57%. |
TunArTTS: Tunisian Arabic Text-To-Speech Corpus (2024.lrec-main)
Copied to clipboard
| Challenge: | Historically, TTS relied on classical methods that proved expensive in terms of data storage and often resulted in robotic-sounding output known as concatenative speech. |
| Approach: | They propose to extract a mono-speaker speech corpus from an online dictionary and use it to develop end-to-end TTS systems for the Tunisian dialect. |
| Outcome: | The proposed system is based on two approaches: training from scratch and transfer learning. |
ImmersiveTTS: Environment-Aware Text-to-Speech with Multimodal Diffusion Transformer and Domain-Specific Representation Alignment (2026.acl-long)
Copied to clipboard
| Challenge: | ImmersiveTTS model synthesizes intelligible speech and environmental audio from natural language descriptions. |
| Approach: | They propose an environment-aware text-to-speech model that integrates natural speech with environmental audio . the model explicitly models cross-modal interactions through a dual-stream stage . |
| Outcome: | Experimental results show that ImmersiveTTS achieves higher naturalness, intelligibility, and audio fidelity than existing approaches. |